Nonparametric Independence Testing for Small Sample Sizes
نویسندگان
چکیده
This paper deals with the problem of nonparametric independence testing, a fundamental decisiontheoretic problem that asks if two arbitrary (possibly multivariate) random variables X,Y are independent or not, a question that comes up in many fields like causality and neuroscience. While quantities like correlation of X,Y only test for (univariate) linear independence, natural alternatives like mutual information of X,Y are hard to estimate due to a serious curse of dimensionality. A recent approach, avoiding both issues, estimates norms of an operator in Reproducing Kernel Hilbert Spaces (RKHSs). Our main contribution is strong empirical evidence that by employing shrunk operators when the sample size is small, one can attain an improvement in power at low false positive rates. We analyze the effects of Stein shrinkage on a popular test statistic called HSIC (Hilbert-Schmidt Independence Criterion). Our observations provide insights into two recently proposed shrinkage estimators, SCOSE and FCOSE we prove that SCOSE is (essentially) the optimal linear shrinkage method for estimating the true operator; however, the nonlinearly shrunk FCOSE usually achieves greater improvements in test power. This work is important for more powerful nonparametric detection of subtle nonlinear dependencies for small samples.
منابع مشابه
A Saddlepoint Approximation to the Limiting Distribution of a K-sample Baumgartner Statistic
Testing hypothesis is one of the most important problems in a nonparametric statistic. Various nonparametric test statistics have been proposed and discussed for a long time. We use the exact critical value for testing hypothesis when the sample sizes are small. However, for large sample sizes, it is very difficult to evaluate the exact critical value. Therefore, the limiting distributions of n...
متن کاملA nonparametric independence test using random permutations
We propose a new nonparametric test for the supposition of independence between two continuous random variables X and Y. Given a sample of (X,Y ), the test is based on the size of the longest increasing subsequence of the permutation which maps the ranks of the X observations to the ranks of the Y observations. We identify the independence assumption between the two continuous variables with th...
متن کاملNonparametric entropy-based tests of independence between stochastic processes
This paper develops nonparametric tests of independence between two stationary stochastic processes. The testing strategy boils down to gauging the closeness between the joint and the product of the marginal stationary densities. For that purpose, I take advantage of a generalized entropic measure so as to build a class of nonparametric tests of independence. Asymptotic normality and local powe...
متن کاملSequential Nonparametric Testing with the Law of the Iterated Logarithm
We propose a new algorithmic framework for sequential hypothesis testing with i.i.d. data, which includes A/B testing, nonparametric two-sample testing, and independence testing as special cases. It is novel in several ways: (a) it takes linear time and constant space to compute on the fly, (b) it has the same power guarantee (up to a small factor) as a nonsequential version of the test with th...
متن کاملNonparametric identification of the minimum effective dose.
We consider identifying the minimum effective dose (MED) in a dose-response study, where the MED is defined to be the lowest dose level producing an effect over that of the zero-dose control. Proposed herein is a nonparametric procedure based on the Mann-Whitney statistic incorporated with the step-down closed testing scheme. A numerical example demonstrates the feasibility of the proposed nonp...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2015